Builds a survey::svydesign() object from a drawn sample, so the analysis
this package does not do — regression, calibration, quantiles with proper
standard errors, replicate weights — can be done by the package that does.
Arguments
- sample
A data frame returned by
draw()withweights = TRUE.- ...
Passed to
survey::svydesign(). Anything named here overrides what the mapping above would have supplied, sonest = TRUEor a replacementfpcis yours to set.
Details
The two packages compute variance from different starting points. This one
uses the design's inclusion probabilities; survey reconstructs the
variance from the design's shape. So the job here is to express each design
in survey's own terms rather than hand over a weight column and hope.
What maps to what
| Design | Expressed as | Standard errors |
design_simple(), design_spatial() | ids = ~1 with fpc the rows the design can reach | identical |
design_reservoir() | ids = ~1 with fpc the rows the stream reached | identical |
design_stratified() | strata from the strata columns, fpc each stratum's size | identical |
design_temporal() | strata from the sampling intervals, fpc each interval's size | identical |
design_cluster() | ids the cluster column, fpc the number of clusters | identical |
design_multistage() | two stages, ids = ~cluster + row, fpc the number of clusters and each cluster's size | identical |
design_weighted(), "poisson" | survey::poisson_sampling(), which models the random size | identical |
design_certainty() | the certainty rows as their own stratum, taken whole | as rest |
design_weighted(), "systematic" | pps = "brewer" with the inclusion probabilities | within about 0.2 percent |
design_spread() | as its first-order design: ids = ~1, or pps = "brewer" with size | survey's is larger |
design_systematic() | ids = ~1 with fpc the frame size | differ |
"Identical" means to floating point, and is checked by this package's tests
against survey::svytotal(), as are domain estimates against
survey::svyby() and degrees of freedom against survey::degf(). The
exceptions are real and worth knowing:
Systematic PPS. Neither package has its joint probabilities. This one uses Deville's approximation and
surveyuses Brewer's; the two are close relatives and agree to a fraction of a percent.Spread.
surveyhas no estimator for a spatially balanced sample, so it is told only the first-order design and computes a variance that ignores the spreading. That errs conservative, often substantially.Systematic.
survey, told onlyids = ~1, computes the simple-random variance;ht_total()uses the successive-difference approximation, which can see a trend along the sort order. Neither is design-unbiased, because no such estimator exists for one systematic sample.
Certainty rows arrive in a stratum where n == N, so survey's own finite
population correction zeroes them out, matching this package's treatment.
Two compositions have no single survey design and are refused rather than
approximated: design_certainty() over a cluster, multistage or Poisson
rest, where the certainty rows and the rest are different kinds of sampling
unit. design_bootstrap() is refused outright — it resamples the sample, so
there is no finite population for svydesign() to represent.
A stratum or interval with a single sampled row makes survey stop with
"Stratum has only one PSU" unless options(survey.lonely.psu) says
otherwise; ht_total() declines in the same situation. Draw at least two
per stratum.
Examples
set.seed(1)
pop <- data.frame(
id = 1:400,
site = rep(c("a", "b", "c", "d"), times = c(200, 100, 60, 40)),
spend = round(stats::runif(400, 10, 500))
)
s <- draw(pop, design_stratified("site", n = 60), seed = 1, weights = TRUE)
des <- as_svydesign(s)
survey::svytotal(~spend, des)
#> total SE
#> spend 108333 6758.6
# The same total, and the same standard error
ht_total(s, "spend")
#> Horvitz-Thompson total (stratified design, n = 60)
#> estimate 108,333.3
#> se 6,758.57 (analytic)
#> 95% CI 94,794.29 to 121,872.4 (t, 56 df)
#> deff 1.02 (about the same as simple random sampling)
# Now the analysis this package does not do
survey::svyquantile(~spend, des, quantiles = 0.5)
#> $spend
#> quantile ci.2.5 ci.97.5 se
#> 0.5 252 198 341 35.69217
#>
#> attr(,"hasci")
#> [1] TRUE
#> attr(,"class")
#> [1] "newsvyquantile"