Skip to contents

Takes every interval-th row, beginning at start. The number of rows drawn follows from the interval and the size of the data, so there is no n.

Usage

design_systematic(interval, start = NULL, order_by = NULL, na_rm = FALSE)

Arguments

interval

Sampling interval. A whole number of 1 or more. Fractional intervals are rejected: they produce uneven gaps, which is not a systematic design.

start

Starting row, in 1:interval. Drawn at random from that range when NULL.

order_by

Optional column to sort by before walking the data. It changes which rows can appear together, so inclusion_prob() and joint_prob() compute against the sorted order too.

na_rm

When order_by is given, drop rows whose sort key is NA instead of raising an error.

Value

A design object, for use with draw().

Variance

A systematic sample has a single random start, so no design-unbiased variance estimator exists. ht_total() uses the successive-difference approximation (Wolter 2007), which compares neighbouring sampled rows in the order the design walked. Sorting on a variable related to what you measure (order_by) is what makes systematic sampling efficient, and this estimator is the one that can see it; see ht_total() for its limits.

References

Madow, W. G. and Madow, L. H. (1944). On the theory of systematic sampling, I. Annals of Mathematical Statistics, 15, 1–24.

Wolter, K. M. (2007). Introduction to Variance Estimation, 2nd ed. Springer.

Examples

df <- data.frame(id = 1:100, value = (1:100) / 10)
draw(df, design_systematic(interval = 10, start = 3))
#>    id value
#> 1   3   0.3
#> 2  13   1.3
#> 3  23   2.3
#> 4  33   3.3
#> 5  43   4.3
#> 6  53   5.3
#> 7  63   6.3
#> 8  73   7.3
#> 9  83   8.3
#> 10 93   9.3
draw(df, design_systematic(interval = 25, order_by = "value"), seed = 1)
#>    id value
#> 1  25   2.5
#> 2  50   5.0
#> 3  75   7.5
#> 4 100  10.0