The probability that a pair of rows both land in the sample. First-order
probabilities from inclusion_prob() give you an unbiased total;
second-order probabilities are what let you put a standard error on it.
Arguments
- data
A data frame.
- design
A design object.
- rows
Optional row indices. Supply these — usually the rows you drew — to get the submatrix for them instead of the full
nrow(data)square, which is what makes this usable on a large population.- simulate
Estimate the probabilities by repeated draws rather than in closed form. Works for every probability design, including the ones with no closed form, at the cost of Monte Carlo error. Refused for
design_bootstrap(), where every row appears in some replicate and the count converges to 1 for all of them.- R
Number of simulated draws when
simulate = TRUE.- seed
Optional seed for the simulation.
Value
A square matrix with one row and column per element of rows
(or per row of data). The diagonal holds first-order probabilities.
Which designs have them
Closed forms exist, and are used, for the designs whose selection is either independent across groups or a simple random sample within them:
design_simple() | n(n-1) / (N(N-1)) for a pair, without replacement |
design_stratified() | within a stratum as above; across strata, independent |
design_cluster() | same cluster: a/A; different clusters: a(a-1)/(A(A-1)) |
design_multistage() | the two stages multiplied, where the per-cluster take is constant |
design_certainty() | 1 between certainty rows; otherwise the other row's own pi |
design_reservoir() | simple random sampling over the first max_items rows |
design_temporal() | within an interval as above; across intervals, independent |
design_spatial() | simple random sampling inside the region |
design_weighted() | "poisson" only, where rows are independent: pi_i * pi_j |
design_systematic() | 1/interval for rows sharing a residue class, otherwise zero |
Systematic sampling is the awkward one. Most pairs can never co-occur, so
their joint probability is genuinely 0 and no design-unbiased variance
estimator exists; ht_total() uses the successive-difference approximation
instead, and says so. Its residue classes follow the order the design walks,
so order_by changes which pairs can co-occur.
Two designs have joint probabilities that are positive but have no closed
form: design_weighted(method = "systematic"), which shuffles the rows
before walking them so every pair can co-occur, and design_spread(). For
both, ht_total() uses a variance approximation that needs only
first-order probabilities, and simulate = TRUE here estimates the joint
probabilities themselves.
Examples
df <- data.frame(id = 1:10)
round(joint_prob(df, design_simple(n = 4)), 3)
#> [,1] [,2] [,3] [,4] [,5] [,6] [,7] [,8] [,9] [,10]
#> [1,] 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#> [2,] 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#> [3,] 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#> [4,] 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133
#> [5,] 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133
#> [6,] 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133
#> [7,] 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133
#> [8,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133
#> [9,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133
#> [10,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400
# Only for the rows you drew
joint_prob(df, design_simple(n = 4), rows = c(2, 5, 7))
#> [,1] [,2] [,3]
#> [1,] 0.4000000 0.1333333 0.1333333
#> [2,] 0.1333333 0.4000000 0.1333333
#> [3,] 0.1333333 0.1333333 0.4000000