Published Updated
SciPy t-tests can return p=0 or p=1 after exact rescaling
SciPy’s one-sample and paired t-tests can reverse a 5% decision after exact power-of-two rescaling because an intermediate variance underflows or overflows.
Current status
We filed SciPy issue #26113 on September 7, 2026 after reproducing the behavior in released versions and recent development code. The issue is open.
On September 13, 2026, maintainer mdhaber posted a run with mpmath set to 30 decimal digits. For both APIs and all three scales, the two-value sample gives t = 3, p ≈ 0.20483276470 and the three-value sample gives t ≈ 5.196152423, p ≈ 0.03509871865. No warnings were captured.
The demonstration uses PR #24840 and mparray, an array backend that computes with mpmath at adjustable precision. Its dtype=float64 label does not mean ordinary NumPy float64 arithmetic. PR #24840 merged into SciPy main on September 23, 2026, providing this high-precision backend path. It does not change the ordinary NumPy float64 calculation reported here, so issue #26113 remains open.
As of September 24, 2026, PR #26135proposes normalizing centered residuals within each reduction slice and restoring scale after the square root. It remains open, mergeable, and unreviewed. Its GitHub Actions jobs are awaiting approval rather than showing a code-test failure. Its author reports regression and high-precision checks; we have not independently rerun the patch. The shared variance implementation is unchanged.
Earlier PR #26114closed without merging. A maintainer requested an AI-use declaration and asked the contributor to await an issue response. PR #26135 includes that declaration.
What we observed
For finite, nonconstant float64 inputs, scipy.stats.ttest_1samp and ttest_rel can returnp=0 or p=1 solely because every observation was multiplied by the same exact power of two. The expected t-statistic and p-value do not change under this positive rescaling.
| Sample | Scale | Expected | SciPy result |
|---|---|---|---|
[s, 2s] | s = 2**-600 | t=3, p≈0.2048 | t=inf, p=0 |
[s, 2s] | s = 2**600 | t=3, p≈0.2048 | t=0, p=1 |
[2s, 3s, 4s] | s = 2**-600 | t=3√3, p≈0.03510 | t=inf, p=0 |
[2s, 3s, 4s] | s = 2**600 | t=3√3, p≈0.03510 | t=0, p=1 |
At a 5% significance threshold, the small two-value sample changes from non-rejection to rejection, while the large three-value sample changes from rejection to non-rejection. This changes the statistical decision in both directions rather than only changing the last digits of a result.
Why the answer should not change
For any positive scale s, the sample [s, 2s] has t-statistic 3 and a two-sided p-value of approximately 0.20483276469913345. The sample [2s, 3s, 4s] has t-statistic 3√3 and a two-sided p-value of approximately 0.03509871864598465. These values follow directly from the sample means, unbiased variances, and standard errors and were also checked with a high-precision Student-t tail calculation independent of SciPy.
At the small scale, the exact sample variances fall below the smallest positivefloat64 value. At the large scale, they exceed the largest finite value. Their square roots and the resulting t-statistics remain representable. The numerical problem is therefore not that the requested final statistic lies outside the available range.
How we checked it
We reproduced the same numerical results with SciPy 1.17.0, 1.17.1, and 1.18.1 release wheels, a 2.0.0.dev0 development wheel, and a source build of SciPy main at commit b881cb1. The tests used NumPy 2.3.5, CPython 3.12, and Linux x86_64.
The inputs are normal finite floating-point numbers with substantial relative variation. Scaling by 2**600 or 2**-600 changes the exponent exactly, so the failure is not caused by decimal conversion or a nearly constant sample. Under NumPy’s default error settings, the small-scale case returns p=0 without a warning. Enabling underflow warnings exposes the intermediate failure.
Reporter-side diagnostic
In the tested development source, ttest_rel delegates to ttest_1samp. The one-sample implementation obtains its variance through _var and _moment, where the centered values are squared before the standard error is formed. On the reported inputs, that intermediate variance becomes 0.0 or inf.
A demonstration that first scales the centered values before computing their norm recovers the expected t-statistics for these examples. This is evidence about the failure mechanism, not a proposed general repair. A production change would also need to address axes, NaN handling, array backends, and related callers of the shared reduction.
This is a reporter-side diagnosis. SciPy has not publicly confirmed the cause or selected a repair strategy.
Scope
The report uses one-sample and paired t-tests as minimal examples because their p-values make the decision change easy to see and their correct values can be derived by hand. NumPy’s variance and standard-deviation reductions show related range loss, and other SciPy functions that use variance reductions can also be affected. The report does not claim that the behavior is exclusive to SciPy or to these two functions.
The scales are deliberate stress tests. This report does not estimate how often ordinary datasets reach the failing range, and it does not propose changing the behavior for samples whose mathematical variance is actually zero.
Relation to verification
This case shows why numerical verification must check transformations that should preserve an answer, not only whether the returned p-value lies between zero and one. Both incorrect p-values look valid in isolation, but an exact rescaling test reveals that the decision is unstable.
The investigation demonstrates the boundary-testing and independent-reference methods behind Licklider’s verification work. It does not mean that nomue currently verifies SciPy, one-sample t-tests, or paired t-tests, and it does not change the current nomue Protocol support boundary.
Public evidence
- SciPy issue #26113— report, reproducer, expected values, tested versions, and current upstream state
- Tested SciPy main source— pinned source for the reporter-side implementation-path diagnosis
- SciPy issue #9562— earlier accepted work avoiding intermediate underflow and overflow in
pearsonr - SciPy issue #20136— open range-protection issue relevant to axis-aware norm calculations