Reports RMSE, MAE, MAPE, SMAPE, \(R^2\) and adjusted \(R^2\) for any
spatial_fit, in one row and on one scale, so that fits from different
backends can be read side by side. Reach for it to score a model on data you
hold out yourself (pass it as newdata), or to get a quick in-sample
reading of how closely a fit tracks its training data.
Usage
model_metrics(object, ...)
# S3 method for class 'spatial_fit'
model_metrics(object, newdata = NULL, ...)Value
A data.frame with n, RMSE, MAE, MAPE, SMAPE, R2, Adj_R2, n_MAPE and
n_SMAPE (the last two are the rows each percentage error was averaged
over; see "Percentage errors on responses with zeros").
Adj_R2 is always NA: GWR's effective parameter count far
exceeds the global predictor count and a GP model has no simple p,
so it is deliberately suppressed. A non-numeric response is an error: a
character or factor response cannot be scored, and used to come back as
n = 0 with every metric NA; a logical response is treated
as 0/1.
Details
It is not a substitute for cross-validation. With newdata = NULL the
numbers are in-sample for a gwr_fit or bayesian_fit, and a
GWR can reach a near-perfect in-sample \(R^2\) at a small bandwidth
without predicting anything. For a figure you can report, use
cv_gwr(), cv_bayes(), cv_rf()
or compare_models_cv().
What the metrics are computed on
With newdata = NULL the metrics come from fitted(object).
That is in-sample for a gwr_fit or a bayesian_fit,
but out-of-bag for an rf_fit, whose fitted() method
returns out-of-bag predictions (see fit_rf_model). The
data.frame model_metrics() returns carries no label distinguishing
the two, so check object$info$fitted_are_oob before comparing
numbers across backends; evaluate_insample() and
compare_models() record it per model in a
metric_basis column. compare_models_cv scores every
backend the same way.
\(R^2\) is \(1 - RSS/TSS\) with the total sum of squares taken about
the mean of the response the model was fitted to. In sample that
is the ordinary \(R^2\). With newdata it is out-of-sample
\(R^2\), the convention every cv_*() function uses: the model is
measured against the prediction it had to beat, the training mean, not
against the new rows' own mean, which it could not have known. It is
below 0 when the model predicts the new rows worse than the training mean
does, and it is NA when the response does not vary about that
baseline by more than rounding error (100 machine epsilons of its
magnitude, whatever its units).
Percentage errors on responses with zeros
MAPE divides by the observed value and SMAPE by
\(|y| + |\hat{y}|\), so neither is defined where its denominator is zero.
Neither returns Inf or NaN. Both are averaged over the rows
whose denominator is non-zero, and are NA when no row qualifies.
Non-zero is judged at the scale of the data: a denominator no larger
than 100 machine epsilons times the largest one counts as zero, so the
rule does not depend on the units of the response.
The n_MAPE and n_SMAPE columns record how many rows that was;
the n column counts finite observation/prediction pairs. Read a
percentage error next to its count: when n_MAPE < n, MAPE is
an average over a subset of the data, whatever its value.
This bites on any response taking exact zeros: counts, rainfall,
abundance, claim amounts. On a zero-inflated response with 62 zeros out of
120, MAPE is an average over the 58 non-zero rows, which
n_MAPE = 58 now says. SMAPE fails differently and more
subtly: it drops the rows where observation and prediction are both near
zero (which on a well-fitted zero-inflated model are the rows it got
right), so it averages the harder rows only and reads worse than
the fit deserves; n_SMAPE shows how many rows it kept, and the
count is only a label, not a repair.
RMSE, MAE and \(R^2\) use every finite row and are
unaffected; prefer them whenever the response can be zero. For a Bayesian
fit, cv_bayes() additionally reports CRPS and interval
coverage, which are proper scoring rules and have no such failure mode.
Which metrics survive a non-Gaussian response
RMSE and MAE are defined for any numeric response and are what to read for a count, a rate or a bounded outcome. MAPE and SMAPE assume a response that is rarely zero (see the previous section), and R-squared and adjusted R-squared compare residual variance to total variance, which is the right comparison for a Gaussian response and a loose one for anything whose variance tracks its mean. None of the four is wrong to compute; each is Gaussian-shaped thinking, and on a Poisson or zero-inflated response should be read as a rough summary rather than a score.
For the Bayesian backend, cv_bayes() additionally reports
CRPS and interval coverage at 50, 80 and 95 percent. Both are proper
scoring rules computed from posterior draws, so they are meaningful for any
family that predicts one number per row (a count, a rate, a binary
or bounded outcome), and they are the numbers to compare when the response
is not Gaussian. A categorical or ordinal family predicts a probability
per response category instead, so cv_bayes() refuses one before
fitting anything. When every fold fails, the
fold_metrics frame cv_bayes() returns carries the CRPS column
but not the coverage_* columns, so code that reads those columns
must tolerate their absence.
See also
Other model evaluation:
compare_models(),
compare_models_cv(),
evaluate_insample(),
residual_morans_i()
Other methods on a fitted model:
coef.bayesian_fit(),
coef.gwr_fit(),
coef.rf_fit(),
fitted.bayesian_fit(),
fitted.gwr_fit(),
fitted.rf_fit(),
predict.bayesian_fit(),
predict.gwr_fit(),
predict.rf_fit(),
print.rf_fit(),
print.spatial_fit(),
residuals.bayesian_fit(),
residuals.gwr_fit(),
residuals.rf_fit(),
summary.spatial_fit()
Examples
if (requireNamespace("ranger", quietly = TRUE)) {
library(sf)
set.seed(1)
pts <- st_as_sf(
data.frame(x = 5e5 + runif(60, 0, 1000), y = 5e6 + runif(60, 0, 1000),
a = rnorm(60)),
coords = c("x", "y"), crs = 32632
)
pts$z <- 2 * pts$a + rnorm(60, 0, 0.3)
# Fit on 40 rows and keep 20 back, so the second call really is out of
# sample; scoring the training rows again would only re-read the fit.
fit <- fit_rf_model(pts[1:40, ], "z", "a", num_trees = 50, seed = 1)
print(model_metrics(fit)) # in-sample (out-of-bag for RF)
model_metrics(fit, newdata = pts[41:60, ]) # on the 20 held-out rows
}
#> n RMSE MAE MAPE SMAPE R2 Adj_R2 n_MAPE n_SMAPE
#> 1 40 0.409009 0.3236812 86.57837 50.87364 0.9541711 NA 40 40
#> n RMSE MAE MAPE SMAPE R2 Adj_R2 n_MAPE n_SMAPE
#> 1 20 0.5080838 0.3836504 158.4334 63.30707 0.9022065 NA 20 20