Beyond “non-significant” results: Why and how to test for practical equivalence

Tomek J., Caldwell A., Eisner DA.

Reporting comparisons with P ≥ 0.05 as showing “no effect” or “no difference” remains one of the most widespread and problematic misinterpretations in the scientific literature. A statistically nonsignificant result shows only that the data do not provide strong evidence for a difference. This distinction matters because such findings can arise for two very different reasons: Either there is no meaningful difference, or a meaningful difference is present but cannot be detected reliably because of limited sample size or high variability. We highlight equivalence testing as a practical framework for distinguishing between these possibilities. Using the two one-sided tests (TOST) approach, investigators can formally test whether effects large enough to be scientifically, clinically, or practically meaningful can be ruled out. We explain the logic of TOST, show how it distinguishes practical equivalence from inconclusive evidence, and illustrate its use in paired and unpaired comparisons. To support broader adoption, we also introduce a freely available online calculator that enables researchers to perform common equivalence tests without coding.

DOI

10.1073/pnas.2611548123

Type

Journal article

Publisher

National Academy of Sciences

Publication Date

2026-08-18T00:00:00+00:00

Volume

123

Permalink More information Close