Published Updated

Checking Welch results with exact rescaling

Exact inputs and independent references expose a changed Welch p-value despite a finite result and no warning in a SciPy boundary test.

A finite answer can still need checking

A six-observation example exposes a numerical failure in SciPy 1.18.1: an exact change of scale leaves the Welch t statistic unchanged, but changes the returned degrees of freedom and p-value. The smaller-scale call emits no warning. We publish the inputs, executable checks and recorded results so another developer can inspect the same behavior.

The useful lesson is a testing method: start with a transformation that should preserve the answer, calculate a reference independently, and inspect each returned field. A successful call and a p-value between zero and one are useful checks, but do not establish numerical accuracy.

Reported to SciPy on September 12, 2026 as issue #26169. Licklider has reproduced the example in the fixed environment below. On September 13, maintainer mdhaber posted the high-precision results described below. The issue remains open.

September 13 update: high-precision results match the reference

In the maintainer’s run, mpmath was set to 100 decimal digits. At scales 1, 2^-300 and 2^300, the reported results all give t ≈ −2.449489743, df = 4 and p ≈ 0.07048399691, with no captured warnings.

The demonstration uses PR #24840 and mparray, an array backend that computes with mpmath at adjustable precision. Its dtype=float64 label does not mean ordinary NumPy float64 arithmetic. As of September 13, 2026, the PR is unmerged and the issue remains open with needs-decision. These posted results do not establish a repair for ordinary NumPy inputs or a released fix.

Since that run, PR #24840 has merged. A separate NumPy float64 repair PR #26209 now proposes normalized Welch calculations for this two-group path and the related Welch ANOVA path. At head 621ef469, all 56 GitHub checks pass and GitHub reports the PR clean and mergeable. The PR and both issues remain open; passing checks do not establish maintainer approval, merge, or release.

Change the scale without changing the problem

Welch’s two-sample t-test compares two independent group means while allowing different variances. Multiplying both groups by the same positive constant preserves its t statistic, degrees of freedom and p-value. The confidence interval should change by that same constant.

Use [-1, 0, 1] and [1, 2, 3]. Each group has sample variance 1 and three observations. The exact reference is t = −√6 and df = 4. Here, df is the degrees-of-freedom parameter used to turn t into a probability.

import numpy as np
from scipy import stats

for k in (0, -300, 300):
    x = np.ldexp(np.array([-1., 0., 1.]), k)
    y = np.ldexp(np.array([1., 2., 3.]), k)
    result = stats.ttest_ind(x, y, equal_var=False)
    print(k, result.statistic, result.df, result.pvalue)

The powers of two in this example preserve the input values exactly. The two-sided reference probability is 1 − (6/5)√(3/5), approximately 0.07048399691.

SciPy 1.18.1 results; the reference df is 4 at every scale
ScaleReturned tReturned dfReturned pCaptured warnings
1−2.4494940.070484None
2^-300−2.4494910.246752None
2^300−2.4494910.246752Three overflow warnings

Both returned probabilities stay above 0.05 in this example. The finding concerns numerical accuracy and the confidence interval, rather than a changed 5% decision. The deliberately extreme scales do not estimate how often this occurs in practical data.

Find the first calculation that loses information

The computed sample variances remain positive and finite: approximately 2.41 × 10^-181 at the small scale and 4.15 × 10^180 at the large scale. The failure occurs when the degrees-of-freedom formula squares the variance contributions. Its numerator and denominator become either both zero or both infinite.

The resulting ratio is NaN, meaning the floating-point calculation did not produce a numerical value. The inspected SciPy 1.18.1 helper replaces NaN degrees of freedom with 1. A separate diagnostic in our reproducer checks this expression and the installed helper. The installed source hash matches that upstream revision.

The helper’s zero-variance explanation does not cover this example: both variances are nonzero. The replacement produces a plausible finite probability using the wrong degrees of freedom. Rescaling the reported 95% interval back to the original units gives approximately [−12.3746, 8.3746], compared with the unscaled call’s [−4.2670, 0.2670].

This differs from our one-sample and paired t-test report, where the variance itself loses range, and from the Welch ANOVA report, where summing group weights overflows. The related SciPy PR #26135 proposes a repair for the one-sample and paired APIs, rather than this Welch helper.

Keep the evidence separate from the implementation

The downloadable check reconstructs each actual binary64 input as an exact rational number and calculates the moments without rounding. A separate script uses mpmath at 100 decimal digits to compare the elementary probability formula with a beta-integral calculation. Neither reference calls SciPy or nomue.

For more complicated cases, interval arithmetic offers another reference form: a narrow range containing the target value. A tolerance check can pass when the whole error range is below its threshold, fail when it is above, and remain unresolved when it crosses the threshold. Such numerical containment concerns the specified calculation; it does not establish that the research assumptions are suitable.

Exact hexadecimal inputs preserve what was actually computed. Saved outputs, warning settings and package versions make the observation repeatable. File hashes identify the evidence, while separate checks establish its numerical meaning. These serve different purposes.

For nomue, an agent-callable verification product, this provides a concrete engineering lesson: test the result delivered at the API boundary, including numerical fields and failure signals. Putting range handling inside a tool reduces dependence on callers adding the right preprocessing. Agents can generate preprocessing, so this is a design motivation, not a claim that they always omit it or that nomue improves model performance.

Download the reproduction instructions, SciPy reproducer, independent reference, recorded outputs and environment, reference output, and file hashes. The run used Python 3.12.14, SciPy 1.18.1, NumPy 2.3.5 and mpmath 1.3.0 on Linux x86_64. It establishes this release-specific example, not a general repair or product ranking.