Takes every interval-th row, beginning at start. The number of rows drawn
follows from the interval and the size of the data, so there is no n.
Arguments
- interval
Sampling interval. A whole number of 1 or more. Fractional intervals are rejected: they produce uneven gaps, which is not a systematic design.
- start
Starting row, in
1:interval. Drawn at random from that range whenNULL.- order_by
Optional column to sort by before walking the data. It changes which rows can appear together, so
inclusion_prob()andjoint_prob()compute against the sorted order too.- na_rm
When
order_byis given, drop rows whose sort key isNAinstead of raising an error.
Value
A design object, for use with draw().
Variance
A systematic sample has a single random start, so no design-unbiased
variance estimator exists. ht_total() uses the successive-difference
approximation (Wolter 2007), which compares neighbouring sampled rows in
the order the design walked. Sorting on a variable related to what you
measure (order_by) is what makes systematic sampling efficient, and this
estimator is the one that can see it; see ht_total() for its limits.
References
Madow, W. G. and Madow, L. H. (1944). On the theory of systematic sampling, I. Annals of Mathematical Statistics, 15, 1–24.
Wolter, K. M. (2007). Introduction to Variance Estimation, 2nd ed. Springer.
Examples
df <- data.frame(id = 1:100, value = (1:100) / 10)
draw(df, design_systematic(interval = 10, start = 3))
#> id value
#> 1 3 0.3
#> 2 13 1.3
#> 3 23 2.3
#> 4 33 3.3
#> 5 43 4.3
#> 6 53 5.3
#> 7 63 6.3
#> 8 73 7.3
#> 9 83 8.3
#> 10 93 9.3
draw(df, design_systematic(interval = 25, order_by = "value"), seed = 1)
#> id value
#> 1 25 2.5
#> 2 50 5.0
#> 3 75 7.5
#> 4 100 10.0