Trains an SVM classifier or regressor using caret (with the kernlab engines), with options for dummy encoding of predictors, feature selection, class-imbalance handling, and hyperparameter tuning via cross-validation. Optional parallel training uses an explicit worker count.
Arguments
- data
A data frame containing predictors and the target.
- target
A string naming the target variable.
- task
Either
"classification"or"regression". Required; there is no default, because guessing it from the target is exactly the mistake this argument exists to prevent.- train_ratio
Training set proportion, strictly between 0 and 1 (default
0.7).- nfolds
Number of CV folds for hyperparameter tuning, a whole number greater than 1 (default
5). Also the number of folds used by the SVM-RFE subset-size search (clamped there to at most the number of training rows);select_method = "rf_rfe"ignores it and uses 10 folds.- kernel
One of
"linear"(default),"radial", or"polynomial".- tune_grid
Optional tuning grid data frame. If
NULL, a default grid for the chosen kernel is used.- feature_select
Logical; if
TRUE, run feature selection on the dummy-encoded training predictors (defaultFALSE).- select_method
Which selector to run when
feature_select = TRUE:"svm_rfe"(default, true SVM-RFE, linear kernel only) or"rf_rfe"(random-forest screening, any kernel). Ignored whenfeature_select = FALSE, and so is the linear-kernel requirement, which is only enforced when SVM-RFE will actually run. An unrecognized value is always an error.- n_features
Optional whole number of features to keep, capped at the number of encoded predictors. For
"svm_rfe"the topn_featuresranked features are kept and the cross-validated size search is skipped; for"rf_rfe"it truncates the selection to its firstn_featuresentries and is the minimum number of features the random-forest fallback keeps (see Details). Ignored whenfeature_select = FALSE. DefaultNULL(the size is chosen automatically).- class_imbalance
Logical; if
TRUEand the task is classification, up-samples classes within CV resampling (defaultFALSE).- seed
Optional whole-number seed (within the R integer range; fractional values are an error, not truncated), applied locally for the duration of the call and restored afterwards; also used to set reproducible RNG streams on parallel workers when
n_cores > 1. DefaultNULL(never seeds by default).- verbose
Logical; if
TRUE, report progress (including each SVM-RFE elimination step). DefaultFALSE.- n_cores
Number of parallel workers (default
1, sequential). Requests are capped at the detected core count; when greater than 1 a cluster is created for the duration of the call and stopped on exit.
Value
An object of class fs_result with:
- selected
Character vector of selected encoded feature names. When
feature_select = FALSEthis is every encoded predictor. With"svm_rfe"it is the surviving subset, ordered from most to least important; with"rf_rfe"it is the subset in the ordercaret::rfe()reports it.- scores
Named numeric vector covering every encoded predictor, not just the survivors: the SVM-RFE criterion (the squared primal weights
w^2of the first, full-feature fit) or, forselect_method = "rf_rfe", the mean random-forest importance recorded across resamples (NAfor predictorsrfe()never scored, or the mean decrease in node impurity when the fallback ran).NULLwhenfeature_select = FALSE, because no selector produced comparable scores.- method
"svm_"followed by the kernel, for example"svm_linear". The selector that ran, if any, is reported indetails$selection$method.- task
"classification"or"regression".- model
The fitted
caret::trainobject.- details
A list with
test_set(the test split with its coerced target and aligned factor levels),predictions(test-set predictions),performance(acaret::confusionMatrixfor classification, or a named RMSE/Rsquared/MAE vector for regression),selection(NULLwhen no selection ran, otherwise a list withmethod,rankingfrom most to least important,scores, and the size search'ssizes,size_scoresandsize_metric– those last three beingNULL,NULLandNAwhenever no size search ran, which is always the case for"rf_rfe"and for"svm_rfe"with an explicitn_features),encoder(the fittedcaret::dummyVarsobject) andn_features(the number of encoded predictors considered, counted before any were dropped).- call
The matched call.
Details
This is the wrapper to reach for when the selector and the final model
should belong to the same family: SVM-RFE ranks features by the weights of
a linear SVM rather than by an external proxy criterion, and the returned
object carries the tuned model and its held-out performance alongside the
chosen features. That comes at a price – a full SVM fit at every
elimination step, plus a cross-validated size search – so on wide data
screen first with a filter such as fs_supervised.
Feature selection (
feature_select = TRUE) defaults to SVM-RFE (Guyon, Weston, Barnhill and Vapnik, 2002). A linear SVM is fitted on the centered and scaled encoded training matrix, the features are ranked by the squared primal weightw^2(recovered from the fit's support vectors and coefficients, summed over the pairwise problems of a multi-class fit), the lowest-ranked feature is dropped (the lowest 10 percent while more than 50 features remain), and the SVM is refitted on the reduced set until one feature is left. Reversing the elimination order gives the ranking, so rank 1 is the feature eliminated last. SVM-RFE requireskernel = "linear", because the ranking criterion is the primal weight vector, which only exists for a linear kernel; combining it with another kernel is an error that points atselect_method = "rf_rfe". The elimination and size-search fits use a fixed cost ofC = 1and are separate from the final model, which is tuned overtune_grid.How many features SVM-RFE keeps: with
n_featuressupplied, exactly that many (the top of the ranking). Otherwise a short ladder of candidate sizes – the powers of two up to the number of features, plus the full size, trimmed to at most six entries but always including 1 and the full size – is scored by nestednfolds-fold cross-validation with the same linear SVM, on folds shared by every candidate size. Inside each fold the standardization and the whole elimination are re-run on that fold's training rows only, and the top features of that fold's own ranking are scored on its held-out rows, so the size search is free of the selection bias of scoring a ranking that has already seen the held-out rows. The winner is the size with the highest mean accuracy (classification) or the lowest mean RMSE (regression), ties going to the smaller size, and the kept features are the top of the ranking computed on the whole training split; if no size could be scored, all features are kept.select_method = "rf_rfe"keeps the older random-forest screening (caret::rfFuncs) and works with every kernel. It runs its own 10-fold cross-validation over subset sizes 1 to p, independent ofnfolds. Ifrfe()fails or selects nothing, a plain random forest is fitted instead, with a warning, and every predictor with a positive mean decrease in node impurity is kept – or, if fewer thann_features(1 whenn_features = NULL) qualify, the topn_featuresby that importance. Withn_featuressupplied the result is then truncated ton_features, so the fallback keeps exactly that many (capped at the number available).Selection always runs on the training split only, so the test set never informs which features survive.
Class-imbalance handling (
class_imbalance = TRUE, classification only) up-samples within each cross-validation resample viacaret::trainControl(sampling = "up"), so no resampled rows leak across folds.After the train/test split, factor (and character) predictor levels in the test set are aligned to the training levels; test rows containing levels unseen in training are dropped with a warning.
A numeric target with
task = "classification"is coerced to a factor (with a message) only when it has at most 10 unique values; more than 10 unique values is an error suggestingtask = "regression".
Suggested packages required at runtime: caret and kernlab
always, e1071 for classification metrics, randomForest when
feature_select = TRUE and select_method = "rf_rfe", and
doParallel/foreach when n_cores > 1.
Examples
# \donttest{
if (requireNamespace("caret", quietly = TRUE) &&
requireNamespace("kernlab", quietly = TRUE) &&
requireNamespace("e1071", quietly = TRUE)) {
res <- fs_svm(
data = iris,
target = "Species",
task = "classification",
nfolds = 3,
kernel = "linear",
tune_grid = data.frame(C = 1),
seed = 42
)
res$details$performance
# SVM-RFE keeps the two most useful measurements
sel <- fs_svm(
data = iris,
target = "Species",
task = "classification",
nfolds = 3,
kernel = "linear",
tune_grid = data.frame(C = 1),
feature_select = TRUE,
select_method = "svm_rfe",
n_features = 2,
seed = 42
)
sel$selected
sel$scores
}
#> Sepal.Length Sepal.Width Petal.Length Petal.Width
#> 0.1598883 0.3103668 4.5281035 5.1138497
# }