Skip to contents

Creates an sf POINT layer of "seed" locations. Multiple strategies are supported: user-provided points, uniform random sampling within a boundary, or k-means clustering of a sampling cloud.

Usage

get_voronoi_seeds(
  boundary = NULL,
  method = c("kmeans", "random", "provided"),
  n = NULL,
  seeds = NULL,
  sample_points = NULL,
  kmeans_nstart = 10,
  kmeans_iter = 100,
  set_seed = NULL
)

Arguments

boundary

Optional polygonal sf object defining the sampling area.

method

One of "kmeans", "random", "provided".

n

Integer; number of seeds to return. Required for method = "kmeans" and method = "random". Ignored for method = "provided", where every row of seeds is returned; a mismatch between n and nrow(seeds) is reported as a warning.

For method = "kmeans" it is an upper bound only: k-means cannot produce more centres than there are distinct positions in the sampling cloud, nor as many centres as there are rows. When n exceeds either ceiling it is clamped, with a warning naming the count actually used. n = nrow(sample_points) is the common case, and yields nrow(sample_points) - 1 seeds. Check nrow() on the result; do not assume n.

Besides a number, n accepts what the level-selection step returned: the integer vector of ranked candidates from determine_optimal_levels() (its first element is used), a select_resolution() result (its $best), or a resolution_profile() (read with select_resolution() at its default criterion). The result then carries attr(, "n_from") saying which.

With method = "kmeans", a selection or profile also carries the centres of the partition the profile scored at that count, and those are returned instead of a new k-means: a k-means partition is the Voronoi partition of its centres, so their cells are the cells the criteria judged. sample_points, when given, is then only assigned to them (for attr(, "kmeans")), and boundary only sets the CRS; kmeans_nstart, kmeans_iter and set_seed are not used. Pass the count as a number (n = sel$best) for a fresh k-means instead.

seeds

sf POINT object of user-provided seeds (method = "provided").

sample_points

Optional sf POINT cloud for k-means clustering. Only the first two coordinate columns are clustered, so a Z or M dimension does not join the distance calculation and dominate it; rows with empty or non-finite coordinates are dropped with a warning, so they never reach stats::kmeans(), which fails on them without naming a cause. A lon/lat cloud, or one with no CRS whose coordinates look like lon/lat (the heuristic ensure_projected() applies, with its warning), is projected before clustering.

kmeans_nstart

Integer; nstart for kmeans(). Default 10. The partition is stats::kmeans(), not the best-of-25 k-means++ run that resolution_profile() scored a count on, so it is not that partition; pass the selection or profile itself as n for that one (see n).

kmeans_iter

Integer; iter.max for kmeans(). Default 100.

set_seed

Optional integer RNG seed.

Value

An sf POINT object with seed_id and method columns. With method = "kmeans" it also carries attr(, "kmeans"), the run behind the seeds: cluster (the seed_id each clustered cloud point was assigned to), rows (those points' positions in the cloud, since rows with unusable coordinates are dropped first), size (points per seed), withinss and tot_withinss (the within-cluster sums of squares), iter and nstart. For the scored centres of a selection or profile (see n) it describes sample_points assigned to them, each point to its nearest seed, with iter NA and nstart the profile's restarts, and is absent when no sample_points were given.

Examples

library(sf)
bnd <- st_sf(geometry = st_sfc(st_polygon(list(rbind(
  c(0, 0), c(100, 0), c(100, 100), c(0, 100), c(0, 0)
))), crs = 32632))
get_voronoi_seeds(bnd, method = "random", n = 5, set_seed = 1)
#> Simple feature collection with 5 features and 2 fields
#> Geometry type: POINT
#> Dimension:     XY
#> Bounding box:  xmin: 20.16819 ymin: 6.178627 xmax: 90.82078 ymax: 94.46753
#> Projected CRS: WGS 84 / UTM zone 32N
#>   seed_id method                  geometry
#> 1       1 random POINT (26.55087 89.83897)
#> 2       2 random POINT (37.21239 94.46753)
#> 3       3 random POINT (57.28534 66.07978)
#> 4       4 random  POINT (90.82078 62.9114)
#> 5       5 random POINT (20.16819 6.178627)