Skip to contents

A design describes how to sample, independently of what you are sampling. Build one with a design_*() constructor, then hand it to draw() with the data. The same design can be reused across data sets, stored, printed and passed around.

The designs

design_simple()

Rows uniformly at random.

design_stratified()

A share of each stratum.

design_systematic()

Every k-th row.

design_cluster()

Whole clusters.

design_multistage()

Clusters, then rows within them.

design_weighted()

Rows with probability governed by a weight.

design_certainty()

Everything above a threshold, plus a sample of the rest.

design_reservoir()

A fixed-size sample from a stream.

design_bootstrap()

Resampled replicates.

design_temporal()

A share of each time interval.

design_spatial()

Rows inside a region.

design_spread()

A sample spread evenly across a map or across auxiliary variables.

What every design guarantees

Arguments mean the same thing everywhere they appear:

  • n is always the total number of rows drawn, never a per-group figure. design_temporal() uses per_interval precisely because that one is per group, and design_cluster() has no n at all because the row count follows from which clusters were selected.

  • allocation always says how a total is split across groups, either "proportional" or "equal". Splitting is exact: the largest-remainder method is used, so a request for 60 rows returns 60 rather than drifting with per-group rounding.

  • na_rm always decides whether rows with a missing key are dropped or raise an error. It is never silently assumed either way.

  • replace always means sampling with replacement within whatever group the design works on, and always rules out draw(weights = TRUE): an inclusion probability describes distinct units, and a sample holding duplicates cannot be weighted by one.

  • draw() always restores the caller's random number stream before returning, and always gives back a data frame with the input's class and column order.

  • Rows come back in frame order for every design that selects a set of rows, including design_systematic() with order_by. The two exceptions are the ones where draw order is meaningful: design_simple() and design_weighted() return rows in the order they were drawn, and design_bootstrap() returns replicates in order with a leading .replicate column.

Estimating from a sample

Because a design is a value, it can report its own inclusion probabilities before any sampling happens. inclusion_prob() gives first-order probabilities and sampling_weight() their reciprocals — the number of population rows each sampled row stands for. joint_prob() gives second-order probabilities, and ht_total() and ht_mean() turn a sample into a population total, mean or proportion — overall or by domain — with a standard error from the variance estimator suited to the design and a confidence interval on the design's degrees of freedom. deff() reports what the design cost in precision against simple random sampling, and sample_summary() reports what was actually drawn against what was in the frame.

Going the other way, plan_size() solves for the sample size a given margin of error requires — the step before choosing a design. For analysis this package does not do, as_svydesign() hands a sample to the survey package.

Not every design has a closed form for these — see inclusion_prob() for which, why, and what to do instead. Those that do not say so rather than returning an approximation.

Choosing one

Every unit equally likely

design_simple(), or design_systematic() when the frame has a useful order, or design_reservoir() when it does not fit in memory.

Guaranteed coverage of subgroups

design_stratified(), with min_per_stratum if rare groups must appear.

Fieldwork cost matters more than efficiency

design_cluster() or design_multistage() — visiting five sites is cheaper than visiting fifty, at the cost of precision.

Large units matter more

design_weighted(), with method = "systematic" or "poisson" if you intend to estimate.

A few units dominate the total

design_certainty() — take those with certainty and sample the tail, which removes them from the variance entirely.

Coverage across time or space

design_temporal() for time; design_spread() for an even spread over a map, which is usually far more precise than a simple random sample when what you measure varies smoothly over space; design_spatial() to restrict sampling to a region.

Balance on known covariates without choosing strata

design_spread() across those covariates.

Uncertainty of a statistic, not a population total

design_bootstrap().