Filters score each feature without fitting a predictive model. They are the cheapest methods in featR and the right first step on wide data. They are also univariate: each feature is judged on its own, so a filter cannot see redundancy or interactions.
The examples below use one simulated data set, so the answers are known in advance:
set.seed(1)
n <- 200
sim <- data.frame(
x_strong = rnorm(n),
x_weak = rnorm(n),
x_noise = rnorm(n),
x_const = 5,
x_gappy = ifelse(runif(n) < 0.6, NA, rnorm(n))
)
sim$x_copy <- sim$x_strong + rnorm(n, sd = 0.05) # near-duplicate
sim$y <- 2 * sim$x_strong + 0.5 * sim$x_weak + rnorm(n)
sim$class <- factor(ifelse(sim$y > 0, "pos", "neg"))
fs_unsupervised(): clean-up without an outcome
fs_unsupervised() scores each column by a target-free
statistic: "variance", "mad",
"iqr", "range", "missing_prop",
or "n_unique". Because it never looks at the outcome, it is
safe to run before a train/test split.
Drop constant columns:
num <- sim[, c("x_strong", "x_weak", "x_noise", "x_const", "x_copy")]
fs_unsupervised(num, method = "variance", threshold = 0)
#> <fs_result> unsupervised_variance
#> Selected 4 of 5 features
#> x_strong, x_weak, x_noise, x_copy
#> Details: mask, indices, filtered, threshold, direction, action, n_features (in $details)Drop columns that are more than half missing. Note
action = "remove". The default action = "keep"
would keep the most missing columns, and featR warns if you
leave all three of threshold, direction, and
action at their defaults for
"missing_prop".
fs_unsupervised(sim[, c("x_strong", "x_gappy")], method = "missing_prop",
threshold = 0.5, action = "remove")$selected
#> [1] "x_strong"Thresholds are on the scale of the statistic: a variance is in
squared units, "missing_prop" is in [0, 1], and
"n_unique" is a count. A threshold chosen for one method
means nothing for another. "mad" uses R’s default
normal-consistency constant (1.4826), so it estimates a standard
deviation.
fs_supervised(): one feature at a time against the
outcome
For a numeric target the score is the absolute Pearson correlation.
For a factor target it is the one-way ANOVA F statistic.
method = "auto" picks the right one.
reg <- sim[, c("x_strong", "x_weak", "x_noise", "x_copy", "y")]
res <- fs_supervised(reg, target = "y", threshold = 0.2)
summary(res)
#> <fs_result> supervised_correlation (regression)
#> Selected 2 of 4 features
#> x_strong, x_copy
#> Details: mask, indices, filtered, threshold, direction, action, n_features (in $details)
#>
#> Call:
#> fs_supervised(data = reg, target = "y", threshold = 0.2)
#>
#> Scores (4 features, ranked):
#> feature score selected
#> x_strong 0.85750 *
#> x_copy 0.85560 *
#> x_weak 0.17000
#> x_noise 0.05669
#> (* = selected)x_copy is kept alongside x_strong even
though it adds nothing. This is the univariate blind spot. The
F scale for a factor target is unbounded, so pick a threshold
on that scale:
cls <- sim[, c("x_strong", "x_weak", "x_noise", "class")]
fs_supervised(cls, target = "class", threshold = 10)$scores
#> x_strong x_weak x_noise
#> 141.5738542 5.4736576 0.2682509Pearson correlation measures linear association
only. A feature with a strong U-shaped effect can score near zero. Use
fs_infogain() or a model-based method if you suspect
non-linear effects.
fs_correlation(): remove redundancy
fs_correlation() ignores the outcome and looks at how
the features relate to each other. With prune = TRUE (the
default) it drops features one at a time until no two survivors
correlate above threshold:
red <- fs_correlation(num[, c("x_strong", "x_weak", "x_noise", "x_copy")],
threshold = 0.9)
red$selected
#> [1] "x_strong" "x_weak" "x_noise"
red$details$dropped
#> [1] "x_copy"
red$details$pairs
#> Var1 Var2 Correlation
#> 1 x_strong x_copy 0.9984364Pruning uses the caret::findCorrelation() heuristic.
Within the strongest remaining pair, it drops the member with the larger
mean absolute correlation to everything else. That choice does not
consult the outcome, so it may drop the member that predicts better.
fs_boruta() prunes by importance instead.
Supported correlation types are "pearson",
"spearman", "kendall",
"polychoric" (ordered factors, needs polycor), and
"pointbiserial" (continuous against dichotomous).
fs_chi(): categorical features against a categorical
outcome
fs_chi() runs a chi-squared test of independence for
every factor or character feature. It applies a multiple-testing
correction (Bonferroni by default) and selects features whose adjusted
p-value is below sig_level. When any expected cell count is
below 5 it switches to a Monte-Carlo p-value, so pass seed
for reproducibility.
cat_data <- data.frame(
colour = factor(ifelse(sim$x_strong > 0, "red", "blue")),
size = factor(sample(c("S", "M", "L"), n, replace = TRUE)),
class = sim$class
)
chi <- fs_chi(cat_data, target = "class", seed = 1)
chi$details$results
#> feature n df p_value adj_p_value significant method
#> 1 colour 200 1 3.090128e-16 6.180256e-16 TRUE asymptotic
#> 2 size 200 2 6.356383e-01 1.000000e+00 FALSE asymptotic
#> correction_applied min_expected
#> 1 TRUE 43.24
#> 2 FALSE 25.76Scores are adjusted p-values, so smaller is
stronger. summary() sorts them ascending. A
p-value measures evidence against independence, not effect size: with
enough rows, a negligible association is still significant. Numeric
columns are ignored, so discretize them first if you want them
tested.
fs_infogain(): bits of uncertainty removed
Information gain is H(Y) - H(Y | X) in bits. Numeric
predictors are discretized into equal-width bins. A numeric target is
discretized once, up front, so that every predictor is scored against
the same classes.
ig <- fs_infogain(sim[, c("x_strong", "x_weak", "x_noise", "class")],
target = "class")
ig$scores
#> x_strong x_weak x_noise
#> 0.41907995 0.11210176 0.05730707Raw information gain favors features with many levels. An ID column
can achieve the maximum possible gain while predicting nothing. Use
normalize = "gain_ratio", which divides by the feature’s
own entropy, when features differ in cardinality:
with_id <- cbind(id = factor(seq_len(n)), sim[, c("x_strong", "class")])
fs_infogain(with_id, "class")$scores
#> id x_strong
#> 0.9953784 0.4190800
fs_infogain(with_id, "class", normalize = "gain_ratio")$scores
#> id x_strong
#> 0.1302194 0.1285049The gain ratio cuts the ID column’s score by roughly a factor of eight, but on this data the ID column still ranks first. The gain ratio reduces the cardinality bias. It does not remove it. Drop identifier and other near-unique columns before you score anything.
By default selected holds every feature with a score
above zero. With finite data almost every feature has some
positive sample gain, so use top_n for a meaningful
cut.