Published
A nonzero Studentized-range tail disappears in SciPy
An exact special-case reference shows SciPy returning zero for a probability near 0.00000002, with no warning in the recorded runs.
A small probability is lost
SciPy returns zero for a Studentized-range upper-tail probability that should be about one in fifty million. The distribution is used in multiple-comparison procedures. An exact formula for a special case gives us a direct way to check this numerical result.
Tasuku Kobayashi submitted the additional reproducer and proposed regression checks to SciPy’s existing distribution-method tracking issue #17832 on September 10, 2026. Upstream confirmation and a correction remain pending.
The input is q=10000, with two groups (k=2) and two degrees of freedom (df=2). The survival function, or SF, gives the probability above the input; the cumulative distribution function, or CDF, gives the probability at or below it.
| Function | SciPy result | Reference |
|---|---|---|
sf(10000, 2, 2) | 0.0 | 1.999999940000002e-8 |
cdf(10000, 2, 2) | 1.0 | 0.9999999800000006 |
Neither call produced a warning in the recorded runs. The expected tail is a normal, nonzero binary64 value, well within the format’s representable range. This example does not reverse a decision at the 5% significance level: both zero and the reference probability are below 0.05.
An exact reference for this case
For two groups and two degrees of freedom, the Studentized-range variable has the same distribution as the absolute value of a Student-t variable with two degrees of freedom, multiplied by the square root of two. For nonnegative q, this gives:
s = sqrt(q*q + 4)
CDF(q) = q / s
SF(q) = 4 / (s * (s + q))The SF expression avoids subtracting nearly equal numbers. The published reference calculation uses integer arithmetic to bound the square root and the resulting probabilities with rational numbers. It checks that both ends of each bound round to the same binary64 value. A separate analytical derivation and rational-arithmetic calculation also checked the reference.
The saved diagnostics show that the internal CDF numerical integration already returns 1.0 before clipping to the interval from zero to one. Clipping is inactive at this input. The final subtraction 1 - cdf therefore cannot be the sole explanation: subtracting the correctly rounded reference CDF would still give a nonzero result, approximately 1.999999943436137e-8.
The detailed cause of the integration error and a general correction remain unresolved. This submission adds a reproducible example to a known numerical problem area; it makes no claim to have discovered a new problem family.
Reproduce the observation
import scipy
from scipy.stats import studentized_range
print("SciPy version:", scipy.__version__)
print("Actual sf: ", studentized_range.sf(10000, 2, 2))
print("Actual cdf:", studentized_range.cdf(10000, 2, 2))The submitted runs cover SciPy 1.18.1 (revision e4e854eaa8f18d807cd3496028e257e36caa93cc) and the official development build 2.0.0.dev0+git20260907.41eeb59 (revision 41eeb590207dc4d8517abd90fc824a3a240832b5). Both used Python 3.12.13 and NumPy 2.3.5 on Linux x86_64.
The relevant class source also matched main at 90d7ff3cf9f2aec2e2840df5cc66dd18cc7dfc2a. That source comparison is separate from executing a build of that revision.
Download the submitted ZIP for the reference, reproducer, proposed tests, outputs, environments, build configurations, and hashes. Keep the Python files together and run:
python reproduce.py
python -m pytest -q test_studentized_range_tail.pyThe proposed regression tests record two passes and two failures on each tested version. The q=100 checks pass; the q=10000 SF and CDF checks fail. These are tests that expose the reported defect, not evidence of a completed fix.
- Code: reference.py, reproduce.py, proposed regression tests.
- Stable release: observed output, test output, environment, build configuration.
- Development build: observed output, test output, environment, build configuration.
- SHA-256 hashes for the eleven submitted files.
What the case demonstrates
A result can lie inside the valid probability interval and still lose a representable tail. An exact special case makes that loss visible and provides a small check that a future implementation can run. This is one way Licklider develops evidence for scientific verification: connect a mathematical reference, executable reproduction, and precisely bounded claim.
Preparation and reference checking used AI assistance, including a separate AI review; the numerical outputs came from Python execution in the recorded environments. The reporter performed the final review and disclosed AI assistance in the upstream submission. This work does not establish a general repair, estimate prevalence in practical datasets, or add Studentized-range verification to nomue’s supported scope.