Fig 1.
Reviewed prevalence of approaches to the interpretation of interaction effects.
The blue plot on the left represents the percentage of approaches used to follow up a statistically significant interaction (N = 221), namely pairwise comparison (N = 206) or descriptive interpretation of the means (N = 15). The green plot on the right further specifies the type of pairwise comparison adopted, namely post-hoc comparisons between all factor levels (N = 204) or a-priori-defined comparisons (N = 4).
Table 1.
Number of studies reviewed reporting different approaches to the interpretation of interaction effects.
Fig 2.
Reviewed prevalence of approaches to the correction for multiple comparisons.
This word-cloud plot represents the frequency of the reported approaches to multiple comparison correction. Font size is scaled by a factor proportional to its frequency so that the bigger, the more frequent. Bonferroni (N = 70), Tukey (N = 53), and Šidák’s (N = 52) corrections were the most adopted. 36 studies reported no correction for multiple comparisons. In all cases, no clear justification for the use of one, multiple, or no type of multiple comparisons correction was provided.
Fig 3.
Two ways to represent and interpret an interaction effect arising from a 2 (group: lesion/control) x 2 (stimuli: fearful/neutral) design.
(A) Barplot with observed means (± standard error) as classically used to represent post-hoc pairwise comparisons between all sublevels involved in an interaction effect (**** = p<0.001, ns = p>0.05). (B) Model estimated marginal means and 95% confidence intervals.
Table 2.
Observed vs estimated means and standard errors.
Fig 4.
Two ways to represent and interpret an interaction effect arising from a three-way Anova having as independent variables: Group (lesion/control), stimulus (fearful/neutral) and working memory ability.
The plot represents the statistically significant group by stimulus interaction. The working memory did not significantly interact with any variable. (A) Barplot with observed means (± standard error) as classically used to represent post-hoc pairwise comparisons between all sublevels involved in an interaction effect (**** = p<0.001, ns = p>0.05). (B) Model estimated marginal means and 95% confidence intervals.
Table 3.
Observed vs estimated means and standard errors.
Fig 5.
Results of 5000 simulations of two-ways Anova with different number of levels (2x2, 2x3, and 3x3), number of subjects per group (N = 5, 10, 20, 30, 50, or 100), and average difference between means (0, 0.1, 0.25, or 0.5).
The null hypothesis was posed as true by looking at differences between observations which have comparable true means, and a standard deviation of 1. The false positive risk represents the relative frequency (%) of either (in orange) Bonferroni-corrected pairwise comparisons based on observed means and errors were statistically significant (p<0.05) or (in green) 95% confidence intervals of the estimated marginal means overlapping for less than 25% of the full CI length.
Fig 6.
Posterior model probabilities of the three models (H1, H2, and H3) were estimated via Bayesian informative hypotheses.
Table 4.
Overview of the characteristics and experimental scenarios in which the approaches to the investigation of the interaction effect discussed din the present paper should or should not be used.