To Remove or not to Remove: the Impact of Outlier Handling on Significance Testing in Testosterone Data
Удалять или не удалять: влияние обработки выбросов на проверку статистической значимости данных о тестостероне
2016-08-29
SCID: 54.1/99uwu6db
Discuss with AI
independent samples t-testoutlier removalrepeated measures ANOVAsignificance testingtestosterone data
Figures from the paper
Abstract (AI)
Outlier removal is common in hormonal research. Here we investigated to what extent removing outliers in hormonal data leads to divergent statistical conclusions. We first show that the most common outlier detection rule is based on a number of standard deviations (SD) from the mean. Next, we used simulations to examine the degree to which statistical conclusions diverge when a test with outlier exclusion yields a statistically significant result whereas the test with outlier inclusion did not, or vice versa (at p = . 05). Simulations were run in duplicate for independent samples t -tests and repeated measures ANOVA designs, and based on real testosterone (T) data and a theoretical gamma distribution of T data. We ran simulations for different sample sizes (30 to 100) and outlier removal rules (2.5 SD and 3 SD). For significant t -tests, we found that in between 14 % to 55 % of the significant cases a test with outlier exclusion yielded a statistically significant result whereas the test with outlier inclusion did not, or vice versa (median p difference: .03–.06). For significant repeated measures ANOVAs, we found that in between 7 % to 28 % of significant cases a test where outlier exclusion yielded a statistically significant result whereas the test with outlier inclusion did not, or vice versa (median p difference: .01–.03). When reporting any test that would lead to a statistically significant result (either the test with inclusion or exclusion of outliers (or both)), in between 5.15 % and 6.89 % of the independent sample t -tests were statistically significant, and for the repeated measures ANOVA design this was between 6.32 % and 7.62 % of the tests. Our results suggest that outlier handling can have a substantial impact on significance testing. We suggest several potential solutions for handling outliers and we argue for a careful assessment of handling outliers in hormonal data.
Key Findings
1
Among significant repeated-measures ANOVAs, 7%–28% produced divergent significance conclusions under outlier exclusion versus inclusion, with median p-value differences of .01–.03.
2
Among significant t-tests, 14%–55% produced divergent significance conclusions depending on whether outliers were excluded or retained, with median p-value differences of .03–.06.
3
Outlier detection in hormonal research most commonly relies on excluding observations beyond a specified number of standard deviations from the mean.
4
Simulations using real and theoretical testosterone data evaluated how 2.5-SD and 3-SD exclusion rules affect independent-samples t-tests and repeated-measures ANOVAs across sample sizes of 30–100.
5
The findings indicate that outlier handling can substantially alter statistical significance in testosterone research, supporting careful assessment and transparent strategies for managing outliers.
Research Object
Testosterone (T) data in hormonal research
Research Subject
The impact of outlier removal rules on statistical significance conclusions in independent-samples t-tests and repeated-measures ANOVA
Publication Details
Publication Date
2016-08-29
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest