A consistent set of wrappers around the analysis of variance designs that
applied work most often needs. Every function takes data first and
names its columns as character strings – six take response and
groups, anova_manova takes responses (plural),
and anova_rm takes response with subject,
within and between. All eight validate their input the same
way and return the same S3 class, anovakit_fit.
Choosing a function
anova_welchOne numeric response, unequal variances permitted. The default choice for a one-way comparison of means.
anova_kwOne numeric response, no distributional assumption. Kruskal-Wallis with Dunn post-hoc. Convert an ordered factor with
as.integer()first.anova_ancovaOne numeric response, adjusting for one or more numeric covariates.
anova_rmOne numeric response measured repeatedly on the same subjects.
anova_manovaTwo or more numeric responses analysed jointly, optionally adjusting for covariates.
anova_binA binary response, via logistic regression.
anova_countA count response, via Poisson or negative binomial regression.
anova_glmAny other generalised linear model family.
Shared conventions
conf_levelsets every interval in the tables the function returns. The one thing it does not reach is$emmeans_object, which is emmeans' own grid and carries emmeans' default level; passlevel =yourself when you summarise it.plots = TRUEbuilds ggplot2 objects and returns them; nothing is ever drawn as a side effect.verbose = FALSEby default; progress goes throughmessage, nevercat. No analysis function writes to the console: warnings and messages raised by car,glmand afex are captured into$notesinstead.Rows with missing or infinite values in the analysed columns are dropped. The exception is the response of
anova_kw: a rank test can use infinite values (as the best or worst possible outcome), so it keeps them.$n_removedcounts the input rows that are not in$data_used, sonrow($data_used) + $n_removedis always the number of rows supplied.Grouping columns are coerced to factors; character columns get their levels in C-locale order, so the reference level does not depend on the session. A factor level that is itself
NAcounts as missing.anova_welchandanova_kwcombine several grouping columns into one factor of the cells that contain data. The model-based functions fit the grouping columns as factors instead. When those factors enter the model additively, marginal means and pairwise comparisons are reported for each factor separately, averaged over the others. When the model contains their interaction, they are reported for the cells: a level combination with no data is left out when the model cannot estimate it, and named in$noteswhen the model estimates it by extrapolation from the other cells.posthoc = FALSEskips the pairwise comparisons and says in$noteshow many there would have been. Above 5000 comparisons they are skipped by default, with a note.Prior weights, where a function accepts them, are named as a column rather than passed as a vector, so they are subsetted with the data. Rows with a weight of zero contribute nothing and are dropped (and counted in
$n_removed). Whole-number weights on a count or 0/1 response are frequency weights: robust standard errors and residual degrees of freedom are then those of the data expanded to one row per observation.vcov_type, where a function accepts it, asks for heteroscedasticity-consistent (sandwich) standard errors. Where the sandwich is degenerate – fitted values on the boundary (separation, a group with no events) or a cell with a single row – the model-based covariance is used instead, with a note.A column may play only one role (response, group, covariate, weights, offset, subject); the names
Residualsand(Intercept)are reserved.Anything the function decided on your behalf, or could not compute, is recorded in
$notes. Read it.
Multiplicity adjustments
adjust takes different values depending on where the comparisons come
from. anova_welch and anova_kw compare directly
and accept the p.adjust methods ("holm",
"hochberg", "hommel", "bonferroni", "BH",
"BY", "fdr", "none"), defaulting to "holm" and
"BH" respectively. The other six go through emmeans and
additionally accept "tukey" (their default), "sidak",
"scheffe" and "dunnettx".
Every $posthoc table reports the unadjusted p-value in
p_value, the adjusted one in p_adjusted and the method in
adjustment. The intervals of anova_welch are
per-comparison intervals; those of the emmeans-based functions are
adjusted (simultaneous) intervals, and a note says when emmeans had
to use a different adjustment for them (it cannot turn a step-down method
such as Holm into intervals, and uses Bonferroni instead). Tukey, Dunnett
and Sidak p-values are computed to a limited precision (ptukey()'s
upper tail is good to about 1e-14), so print() shows those below
1e-10 (1e-12 for Sidak) as a bound; the stored values are unchanged.
See also
anovakit_fit for the returned object.
Author
Maintainer: Justin Chase jchase.msu@gmail.com