Skip to contents

The probability that a pair of rows both land in the sample. First-order probabilities from inclusion_prob() give you an unbiased total; second-order probabilities are what let you put a standard error on it.

Usage

joint_prob(data, design, rows = NULL, simulate = FALSE, R = 5000, seed = NULL)

Arguments

data

A data frame.

design

A design object.

rows

Optional row indices. Supply these — usually the rows you drew — to get the submatrix for them instead of the full nrow(data) square, which is what makes this usable on a large population.

simulate

Estimate the probabilities by repeated draws rather than in closed form. Works for every probability design, including the ones with no closed form, at the cost of Monte Carlo error. Refused for design_bootstrap(), where every row appears in some replicate and the count converges to 1 for all of them.

R

Number of simulated draws when simulate = TRUE.

seed

Optional seed for the simulation.

Value

A square matrix with one row and column per element of rows (or per row of data). The diagonal holds first-order probabilities.

Which designs have them

Closed forms exist, and are used, for the designs whose selection is either independent across groups or a simple random sample within them:

design_simple()n(n-1) / (N(N-1)) for a pair, without replacement
design_stratified()within a stratum as above; across strata, independent
design_cluster()same cluster: a/A; different clusters: a(a-1)/(A(A-1))
design_multistage()the two stages multiplied, where the per-cluster take is constant
design_certainty()1 between certainty rows; otherwise the other row's own pi
design_reservoir()simple random sampling over the first max_items rows
design_temporal()within an interval as above; across intervals, independent
design_spatial()simple random sampling inside the region
design_weighted()"poisson" only, where rows are independent: pi_i * pi_j
design_systematic()1/interval for rows sharing a residue class, otherwise zero

Systematic sampling is the awkward one. Most pairs can never co-occur, so their joint probability is genuinely 0 and no design-unbiased variance estimator exists; ht_total() uses the successive-difference approximation instead, and says so. Its residue classes follow the order the design walks, so order_by changes which pairs can co-occur.

Two designs have joint probabilities that are positive but have no closed form: design_weighted(method = "systematic"), which shuffles the rows before walking them so every pair can co-occur, and design_spread(). For both, ht_total() uses a variance approximation that needs only first-order probabilities, and simulate = TRUE here estimates the joint probabilities themselves.

Examples

df <- data.frame(id = 1:10)
round(joint_prob(df, design_simple(n = 4)), 3)
#>        [,1]  [,2]  [,3]  [,4]  [,5]  [,6]  [,7]  [,8]  [,9] [,10]
#>  [1,] 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#>  [2,] 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#>  [3,] 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133 0.133
#>  [4,] 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133 0.133
#>  [5,] 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133 0.133
#>  [6,] 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133 0.133
#>  [7,] 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133 0.133
#>  [8,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133 0.133
#>  [9,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400 0.133
#> [10,] 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.133 0.400

# Only for the rows you drew
joint_prob(df, design_simple(n = 4), rows = c(2, 5, 7))
#>           [,1]      [,2]      [,3]
#> [1,] 0.4000000 0.1333333 0.1333333
#> [2,] 0.1333333 0.4000000 0.1333333
#> [3,] 0.1333333 0.1333333 0.4000000