Generates bootstrap replicates. "simple" resamples rows independently;
"block" is a moving-block bootstrap, which preserves the serial dependence
in ordered data by concatenating ceiling(n / block_length) independently
chosen blocks per replicate.
Usage
design_bootstrap(
n_replicates = 1000L,
n = NULL,
method = c("simple", "block"),
block_length = NULL
)Arguments
- n_replicates
Number of replicates to generate.
- n
Rows per replicate.
NULLusesnrow(data), the standard nonparametric bootstrap.- method
"simple"or"block".- block_length
Block length for
method = "block", and an error formethod = "simple", where it would have no effect.NULLusesfloor(nrow(data)^(1/3)), a common rule of thumb.
Value
A design object, for use with draw().
Details
draw() returns all replicates in one data frame with a leading .replicate
column. Split them with split(out, out$.replicate).
Not a survey design
The bootstrap resamples the data you already have to approximate the
sampling distribution of a statistic (Efron 1979). It is not a probability
sample of a finite population, so it has no inclusion probabilities and
cannot be passed to ht_total(). For the design-based variance of a total
or mean, use the estimators there; for a statistic they do not cover, such
as a median, this is the tool. Blocks preserve short-range dependence in
ordered data (Künsch 1989).
References
Efron, B. (1979). Bootstrap methods: another look at the jackknife. Annals of Statistics, 7, 1–26.
Künsch, H. R. (1989). The jackknife and the bootstrap for general stationary observations. Annals of Statistics, 17, 1217–1241.
Examples
df <- data.frame(id = 1:100, value = (1:100) / 10)
reps <- draw(df, design_bootstrap(n_replicates = 5, n = 20), seed = 1)
table(reps$.replicate)
#>
#> 1 2 3 4 5
#> 20 20 20 20 20
# A statistic per replicate
vapply(split(reps, reps$.replicate), function(r) mean(r$value), numeric(1))
#> 1 2 3 4 5
#> 5.450 5.285 4.395 4.735 5.570