A design describes how to sample, independently of what you are sampling.
Build one with a design_*() constructor, then hand it to draw() with the
data. The same design can be reused across data sets, stored, printed and
passed around.
The designs
design_simple()Rows uniformly at random.
design_stratified()A share of each stratum.
design_systematic()Every k-th row.
design_cluster()Whole clusters.
design_multistage()Clusters, then rows within them.
design_weighted()Rows with probability governed by a weight.
design_certainty()Everything above a threshold, plus a sample of the rest.
design_reservoir()A fixed-size sample from a stream.
design_bootstrap()Resampled replicates.
design_temporal()A share of each time interval.
design_spatial()Rows inside a region.
design_spread()A sample spread evenly across a map or across auxiliary variables.
What every design guarantees
Arguments mean the same thing everywhere they appear:
nis always the total number of rows drawn, never a per-group figure.design_temporal()usesper_intervalprecisely because that one is per group, anddesign_cluster()has nonat all because the row count follows from which clusters were selected.allocationalways says how a total is split across groups, either"proportional"or"equal". Splitting is exact: the largest-remainder method is used, so a request for 60 rows returns 60 rather than drifting with per-group rounding.na_rmalways decides whether rows with a missing key are dropped or raise an error. It is never silently assumed either way.replacealways means sampling with replacement within whatever group the design works on, and always rules outdraw(weights = TRUE): an inclusion probability describes distinct units, and a sample holding duplicates cannot be weighted by one.draw()always restores the caller's random number stream before returning, and always gives back a data frame with the input's class and column order.Rows come back in frame order for every design that selects a set of rows, including
design_systematic()withorder_by. The two exceptions are the ones where draw order is meaningful:design_simple()anddesign_weighted()return rows in the order they were drawn, anddesign_bootstrap()returns replicates in order with a leading.replicatecolumn.
Estimating from a sample
Because a design is a value, it can report its own inclusion probabilities
before any sampling happens. inclusion_prob() gives first-order
probabilities and sampling_weight() their reciprocals — the number of
population rows each sampled row stands for. joint_prob() gives
second-order probabilities, and ht_total() and ht_mean() turn a sample
into a population total, mean or proportion — overall or by domain — with a
standard error from the variance estimator suited to the design and a
confidence interval on the design's degrees of freedom. deff() reports
what the design cost in precision against simple random sampling, and
sample_summary() reports what was actually drawn against what was in the
frame.
Going the other way, plan_size() solves for the sample size a given margin
of error requires — the step before choosing a design. For analysis this
package does not do, as_svydesign() hands a sample to the survey package.
Not every design has a closed form for these — see inclusion_prob() for
which, why, and what to do instead. Those that do not say so rather than
returning an approximation.
Choosing one
- Every unit equally likely
design_simple(), ordesign_systematic()when the frame has a useful order, ordesign_reservoir()when it does not fit in memory.- Guaranteed coverage of subgroups
design_stratified(), withmin_per_stratumif rare groups must appear.- Fieldwork cost matters more than efficiency
design_cluster()ordesign_multistage()— visiting five sites is cheaper than visiting fifty, at the cost of precision.- Large units matter more
design_weighted(), withmethod = "systematic"or"poisson"if you intend to estimate.- A few units dominate the total
design_certainty()— take those with certainty and sample the tail, which removes them from the variance entirely.- Coverage across time or space
design_temporal()for time;design_spread()for an even spread over a map, which is usually far more precise than a simple random sample when what you measure varies smoothly over space;design_spatial()to restrict sampling to a region.- Balance on known covariates without choosing strata
design_spread()across those covariates.- Uncertainty of a statistic, not a population total