Skip to contents

Selects clusters at the first stage, then draws rows within them at the second. n is the total across the selected clusters under both allocations, matching design_stratified().

Usage

design_multistage(
  clusters,
  n_clusters,
  n,
  allocation = c("equal", "proportional"),
  min_per_cluster = 0L,
  replace = FALSE,
  na_rm = FALSE
)

Arguments

clusters

Column naming each row's cluster.

n_clusters

Number of clusters to select at stage one.

n

Total rows to draw across the selected clusters. Make it a multiple of n_clusters, and no larger than n_clusters times the smallest cluster, if you intend to estimate: otherwise the per-cluster take depends on which clusters were selected and there is no closed-form inclusion probability. See inclusion_prob().

allocation

"equal" splits n evenly across the selected clusters; "proportional" splits it in proportion to their size.

min_per_cluster

Minimum rows from each selected cluster. Defaults to 0. Take at least two rows per cluster if you intend to estimate a variance: with one, the variation within clusters cannot be measured and ht_total() falls back to the ultimate-cluster jackknife.

replace

Sample with replacement within each cluster?

na_rm

Drop rows whose cluster label is NA instead of raising an error.

Value

A design object, for use with draw().

Examples

df <- data.frame(id = 1:100, site = rep(paste0("s", 1:10), each = 10))
nrow(draw(df, design_multistage("site", n_clusters = 4, n = 12), seed = 1))
#> [1] 12