Ensures point geometry, projected CRS, and removes rows with missing or
non-finite values in modeling columns, including rows whose
geometry is empty or whose coordinates are not finite, which no model
backend can use. All non-POINT geometries (including MULTIPOINT) are
coerced to representative points via coerce_to_points(), so
downstream coordinate extraction always aligns one row per observation.
Any Z or M coordinate (POINT Z from a GPS, a GeoPackage or KML) is
dropped, because every backend works in 2-D map distance.
Usage
prep_model_data(
data_sf,
response_var,
predictor_vars,
boundary = NULL,
pointize = c("auto", "surface", "point_on_surface", "centroid", "line_midpoint",
"bbox_center"),
require_response = TRUE
)Arguments
- data_sf
An sf object.
- response_var
Response variable column name.
- predictor_vars
Predictor column names. May be
character(0)for an intercept-only model (fit_bayesian_spatial_model()supports one; the GWR and random-forest backends do not and reject it themselves).- boundary
Optional sf/sfc for CRS alignment.
- pointize
Strategy for non-point geometry coercion, passed to
coerce_to_points. One of"auto"(default),"centroid","point_on_surface","surface"(an alias for"point_on_surface"),"line_midpoint"or"bbox_center".- require_response
Logical; if FALSE the response column is not required to be present (useful for out-of-sample prediction where the response is unknown). Default TRUE.
Value
An sf object with 2-D (XY) POINT geometry, cleaned of rows
carrying missing or non-finite values in the modelling columns or in the
coordinates. What
was removed is recorded on the attribute "dropped", a list with
n (rows dropped), n_geometry (how many of them for an
empty or non-finite geometry), which (their positions in
data_sf), row_id (their ..row_id values when the
layer carries that column, else NULL) and reason (one per
dropped row: "geometry", "missing" or
"non_finite", in that order of precedence when several apply),
plus n_rows, the number of rows returned, which the record was
made for.
Every fit stores n as $info$n_dropped. The record
describes the rows this call returned and does not survive subsetting:
clean[i, ], like dplyr::filter(), slice() or
arrange() of it, is a plain layer with no "dropped"
attribute, and a fit given such a subset with .already_prepped =
TRUE reports n_dropped = 0 even when the parent layer dropped
rows. sf::st_drop_geometry() keeps the record, since the rows are
the same; see [.spatialkit_rows for what binding such data
frames does. The CRS is
projected whenever one can be established. A CRS-less layer is decided by
the lon/lat heuristic (see ensure_projected): if its bounding
box fits the lon/lat envelope and it either spans more than one unit
on some axis or carries decimal-degree-like precision, it is read as
EPSG:4326 and projected, with a warning. A small planar survey inside that
envelope is included in that, deliberately; only coordinates the heuristic
declines are passed through as-is. Set the CRS on data_sf if the
data are planar.
Details
The response may not appear in predictor_vars. Using it as its own
predictor is leakage no backend catches: an out-of-bag R^2 near 1 in the
random forest, a silently reduced design matrix in GWR, duplicated rows and
a phantom <none> entry in the GWR selection table. It is refused
here.
Column names must be syntactically valid R names (make.names(x) == x).
Every backend builds a model formula from these names. A name R parses as
an expression would fit a different model from the one requested while the
fit object still recorded the name you asked for: "B5-B4" is
B5 - B4, and "log(a)" is a function call. Rename the column
(for example with make.names()) before fitting.
See also
Other model fitting:
fit_bayesian_spatial_model(),
fit_gwr_model(),
fit_rf_model(),
gp_lengthscale_bounds(),
new_spatial_fit()
Other spatial data preparation:
clip_target_for(),
coerce_to_points(),
ensure_projected(),
harmonize_crs()
Examples
library(sf)
dat <- st_as_sf(
data.frame(x = 1:5, y = 5:1,
resp = c(1, 2, NA, 4, 5),
pred = c(1, 2, 3, 4, Inf)),
coords = c("x", "y"), crs = 32632
)
clean <- prep_model_data(dat, "resp", "pred") # drops rows 3 (NA) and 5 (Inf)
attr(clean, "dropped")
#> $n
#> [1] 2
#>
#> $n_geometry
#> [1] 0
#>
#> $which
#> [1] 3 5
#>
#> $row_id
#> NULL
#>
#> $reason
#> [1] "missing" "non_finite"
#>
#> $n_rows
#> [1] 3
#>