Selects clusters at the first stage, then draws rows within them at the
second. n is the total across the selected clusters under both allocations,
matching design_stratified().
Usage
design_multistage(
clusters,
n_clusters,
n,
allocation = c("equal", "proportional"),
min_per_cluster = 0L,
replace = FALSE,
na_rm = FALSE
)Arguments
- clusters
Column naming each row's cluster.
- n_clusters
Number of clusters to select at stage one.
- n
Total rows to draw across the selected clusters. Make it a multiple of
n_clusters, and no larger thann_clusterstimes the smallest cluster, if you intend to estimate: otherwise the per-cluster take depends on which clusters were selected and there is no closed-form inclusion probability. Seeinclusion_prob().- allocation
"equal"splitsnevenly across the selected clusters;"proportional"splits it in proportion to their size.- min_per_cluster
Minimum rows from each selected cluster. Defaults to
0. Take at least two rows per cluster if you intend to estimate a variance: with one, the variation within clusters cannot be measured andht_total()falls back to the ultimate-cluster jackknife.- replace
Sample with replacement within each cluster?
- na_rm
Drop rows whose cluster label is
NAinstead of raising an error.
Value
A design object, for use with draw().
Examples
df <- data.frame(id = 1:100, site = rep(paste0("s", 1:10), each = 10))
nrow(draw(df, design_multistage("site", n_clusters = 4, n = 12), seed = 1))
#> [1] 12