Researchers from the Universities of Oxford, Manchester and Arkansas are urging scientists to stop treating “non-significant” results as proof that nothing happened.
The common statistical mistake, they warn in a new paper in PNAS - could be leading researchers to draw the wrong conclusions from their data.
A p-value, which measures how surprising the observed data would be if there were truly no effect, of greater than 0.05 does not, they say, show there is “no effect”. However, the interpretation remains widespread in about 50% of research papers and conference presentations, according to sources.
Instead, the authors say, a non-significant result simply means there is insufficient evidence to conclude a difference exists, not proof that a difference does not exist. Crucially, the same statistical result can arise either because there is genuinely no meaningful effect or because an important effect is hidden by small sample sizes or highly variable data.
The researchers argue that failing to recognise this distinction risks oversimplifying scientific findings and may cause potentially important effects to be overlooked. With many studies not having a large-enough sample size to detect the real difference, reporting that there is none can miss promising therapies or fail to detect genuine risk, derailing subsequent research in the area.
Equivalence testing, the paper argues, guards against declaring an effect unimportant when the data cannot support that claim. A small or noisy study will usually return an inconclusive result rather than a verdict of no difference, which is the honest answer.
To address the problem, the team is promoting a statistical approach known as equivalence testing. Rather than asking whether there is evidence for a difference, equivalence testing asks whether any difference that exists is too small to matter in scientific, clinical or practical terms.
DPAG’s Jakob Tomek, a co-author on this paper explains, ‘The great thing about equivalence testing is that it relies on statistical concepts everyone knows and is entirely compatible with mainstream statistics. And through the online calculator we developed, anyone can start using it right away with ease to draw better conclusions from their data.’
The method allows researchers to distinguish between effects that are genuinely negligible and results that remain inconclusive because there is not enough reliable evidence.
The research focussed on the two one-sided tests procedure, or TOST, which has already gained traction in psychology, medicine and pharmaceutical regulation but remains underused across many areas of the life and natural sciences.
Wider adoption of equivalence testing, the researcher add, could improve the quality of scientific reporting and help prevent non-significant findings from being misrepresented as evidence of no effect.
Co-author, David Eisner, Professor of Cardiac Physiology from The University of Manchester, said, ‘Researchers are often interested in whether an effect is absent or too small to be important, but traditional statistical testing cannot answer that question.
A non-significant result is frequently interpreted as proof that nothing happened, when it may simply mean there is not enough evidence to be certain.’
Jakub Tomek comments ‘Equivalence testing helps separate genuinely trivial effects from unresolved questions, giving scientists a much clearer picture of what their data are actually telling them and improving confidence in the conclusions that are reported.
We hope the method will encourage more careful interpretation of scientific data and improve the way research findings are reported, understood and acted upon.’
To make the approach more accessible, the team has also developed a free online calculator that allows researchers to perform common equivalence tests without writing computer code.
The paper is out now: https://www.pnas.org/doi/10.1073/pnas.2611548123

