Forms the Horvitz-Thompson total sum(y / pi) with a standard error, a
confidence interval and a design effect, using a variance estimator matched
to how the design actually randomised.
Usage
ht_total(
sample,
y,
variance = c("auto", "analytic", "jackknife", "none"),
level = 0.95,
by = NULL,
df = NULL
)Arguments
- sample
A data frame returned by
draw()withweights = TRUE.- y
The variable to total: a column name, or a numeric vector as long as
sample.- variance
How to compute it.
"auto"uses the design's own estimator (see "Variance") and falls back to the jackknife when that declines;"analytic"insists on the design's own estimator, returningNAwith the reason innoterather than falling back;"jackknife"always resamples;"none"skips it. The result reports which was used inmethod, and"none"when nothing could produce a figure.- level
Confidence level for the interval.
- by
Optional column name(s) in
sampledefining domains. When given, the result is one row per domain. See "Domains".- df
Degrees of freedom for the interval.
NULLuses the design's (see "Confidence intervals");Infgives a normal interval.
Value
A list with a print() method, holding:
totalThe Horvitz-Thompson total,
sum(y / pi).variance,se,ci,level,dfIts estimated variance, standard error, confidence interval, and the degrees of freedom the interval used.
NAwhere the design supports no variance.n,designRows used, and the design's type.
deffThe design effect — see
deff().methodWhich estimator produced the variance. See "Variance".
noteWhy a variance is missing, what approximation was used, or which fallback was taken.
NULLwhen an exact estimator applied cleanly.
With by, a data frame of class drawn_by instead: one row per domain,
with the domain columns, n, total, se, ci_lower, ci_upper and
method.
Variance
Each design gets the estimator the survey-sampling literature recommends for
it, and method in the result says which one was used:
"analytic"Exact and design-unbiased. Fixed-size designs with known joint probabilities use the Sen-Yates-Grundy estimator; Poisson sampling, whose rows are independent, uses
sum((1 - pi) / pi^2 * y^2); cluster designs apply the estimator at the cluster level, because taking whole clusters makes the row count random; certainty designs hand the problem torest, since certainty rows are in every sample and add nothing."deville"Systematic probability-proportional-to-size (
design_weighted()withmethod = "systematic"). Its joint probabilities have no closed form, so this uses Deville's (1999) approximation, which needs only the first-order probabilities and is the best-performing of the standard approximations for high-entropy designs (Matei and Tillé 2005). Randomised systematic selection is close to high entropy (Hartley and Rao 1962). Unlike a plain jackknife, each row's contribution is scaled by its own1 - pi, so rows near certainty stop adding variance they do not have."successive difference"design_systematic(). No design-unbiased estimator exists — the design has one random start, and most pairs of rows can never be drawn together — so this uses the successive-difference approximation, which compares each sampled row with its neighbour in the order the design walked (Wolter 2007). It removes a trend along that order rather than counting it as noise, which is why it beats treating the sample as simple random: on a frame sorted by a variable related toy, the simple random formula overstated the variance two- to thirtyfold in the package's simulations, where this one stayed near the truth. Two limits no estimator from a single systematic sample can escape: it is understated if the frame cycles with a period matchinginterval, and when a smooth trend so dominatesythat the random start is almost the only source of variation, it runs low (0.3 to 0.7 of the truth when the trend spanned 25 noise standard deviations)."local mean"design_spread(). A spatially balanced sample is well spread precisely because neighbours are rarely taken together, so this compares each sampled row with the local mean of its nearest sampled neighbours (Stevens and Olsen 2003; Grafström and Schelin 2014)."jackknife"The delete-one-unit jackknife, by stratum (JKn) where the design has strata, and deleting whole clusters where it has them. Used when requested, and as the fallback under
variance = "auto".
The analytic estimator declines — returning NA with the reason in note
— rather than understate. Two cases matter in practice:
A stratum or interval with a single sampled row. There is no variation inside it to measure, and simply leaving it out, as the Sen-Yates-Grundy sum does, reports a standard error that is too small — zero, if every stratum is like that. Draw at least two per stratum (
min_per_stratum = 2,per_interval = 2).One row per cluster in a multistage design. The within-cluster variation is unmeasurable and the exact estimator understates by around a third.
variance = "auto"falls back to the delete-a-cluster jackknife, which is the ultimate-cluster approximation standard in survey practice.
A Sen-Yates-Grundy estimate that comes out negative is treated the same way: a failure of the estimator on an unlucky sample, not a variance.
Confidence intervals
Intervals use Student's t with the design's degrees of freedom — the number
of primary sampling units minus the number of strata — rather than the
normal distribution (Korn and Graubard 1999). The difference is negligible
with hundreds of units and decisive with a handful: at four clusters a
normal interval covers about 87% where it claims 95%, and the t interval
restores most of that. The degrees of freedom used are reported as df;
pass df = Inf for the normal interval.
Domains
by estimates for each level of one or more columns — sites, regions,
months — without subsetting the sample first. Subsetting a sample and
treating the piece as its own sample gets the variance wrong, because the
number of rows that happen to fall in a domain is itself random. Here each
domain's total is estimated from y * (row in domain) over the whole sample,
so the same design and the same variance machinery apply (Särndal, Swensson
and Wretman 1992, section 10.3).
References
Horvitz, D. G. and Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47, 663–685.
Sen, A. R. (1953). On the estimate of the variance in sampling with varying probabilities. Journal of the Indian Society of Agricultural Statistics, 5, 119–127.
Yates, F. and Grundy, P. M. (1953). Selection without replacement from within strata with probability proportional to size. Journal of the Royal Statistical Society B, 15, 253–261.
Hartley, H. O. and Rao, J. N. K. (1962). Sampling with unequal probabilities and without replacement. Annals of Mathematical Statistics, 33, 350–374.
Deville, J.-C. (1999). Variance estimation for complex statistics and estimators: linearization and residual techniques. Survey Methodology, 25, 193–203.
Matei, A. and Tillé, Y. (2005). Evaluation of variance approximations and estimators in maximum entropy sampling with unequal probability and fixed sample size. Journal of Official Statistics, 21, 543–570.
Wolter, K. M. (2007). Introduction to Variance Estimation, 2nd ed. Springer.
Korn, E. L. and Graubard, B. I. (1999). Analysis of Health Surveys. Wiley.
Särndal, C.-E., Swensson, B. and Wretman, J. (1992). Model Assisted Survey Sampling. Springer.
Examples
set.seed(1)
pop <- data.frame(
id = 1:200,
site = rep(c("a", "b"), times = c(150, 50)),
spend = round(stats::runif(200, 10, 500))
)
s <- draw(pop, design_stratified("site", n = 40), seed = 1, weights = TRUE)
ht_total(s, "spend")
#> Horvitz-Thompson total (stratified design, n = 40)
#> estimate 54,000
#> se 3,713.698 (analytic)
#> 95% CI 46,482.01 to 61,517.99 (t, 38 df)
#> deff 1.03 (about the same as simple random sampling)
sum(pop$spend) # the truth
#> [1] 52732
# One estimate per site, from the same sample
ht_total(s, "spend", by = "site")
#> Horvitz-Thompson total by domain (stratified design, 95% CI, t, 38 df)
#> site n total se ci_lower ci_upper method
#> a 30 40715 3152 34334 47096 analytic
#> b 10 13285 1963 9310 17260 analytic
tapply(pop$spend, pop$site, sum)
#> a b
#> 39123 13609