Published Updated

Exact rescaling can reverse SciPy’s Welch ANOVA decision

At an extreme input scale, SciPy’s Welch ANOVA changes a p-value from 0.02650 to 0.05611, crossing the 5% threshold without losing input information.

The same data, a different decision

Multiplying every observation by the same positive constant should preserve Welch’s ANOVA result. This test compares the means of several independent groups while allowing their variances to differ.

In a deliberately extreme example using SciPy 1.18.1, scaling every observation by 2**-511 changes the returned p-value from approximately 0.02650 to 0.05611. The first is below the conventional 5% significance threshold; the second is above it. The scaled observations retain the original input information exactly.

Tasuku Kobayashi reported the example in SciPy issue #26146 on September 10, 2026. On September 13, maintainer mdhaber posted the high-precision results described below. The issue remains open.

September 13 update: high-precision results preserve the decision

In the maintainer’s run, mpmath was set to 30 decimal digits. Both the original and scaled samples return F ≈ 10.28571428571 and p ≈ 0.02650081125, matching the reference and staying below 0.05.

The demonstration uses PR #24840 and mparray, an array backend that computes with mpmath at adjustable precision. Its dtype=float64 label does not mean ordinary NumPy float64 arithmetic. As of September 13, 2026, the PR is unmerged and the issue remains open with needs-decision. These posted results do not establish a repair for ordinary NumPy inputs or a released fix.

Since that run, PR #24840 has merged. A separate NumPy float64 repair PR #26209 now proposes normalized weights for this Welch ANOVA path and a normalized degrees-of-freedom calculation for the related two-group Welch path. At head 621ef469, all 56 GitHub checks pass and GitHub reports the PR clean and mergeable. The PR and both issues remain open; passing checks do not establish maintainer approval, merge, or release.

A nine-observation example

The report uses three groups: [-1, 0, 1], [1, 2, 3], and [3, 4, 5]. Both calls explicitly select Welch’s ANOVA with equal_var=False.

import numpy as np
from scipy import stats

a = np.array([-1, 0, 1], dtype=np.float64)
b = np.array([1, 2, 3], dtype=np.float64)
c = np.array([3, 4, 5], dtype=np.float64)
scale = 2.0 ** (-511)

print("Original:", stats.f_oneway(a, b, c, equal_var=False))
print("Scaled:  ", stats.f_oneway(a * scale, b * scale, c * scale, equal_var=False))
Reference and reported SciPy results
CalculationF-statisticp-valueBelow 0.05?
Welch reference, either scale72/7 ≈ 10.28571449/1849 ≈ 0.02650081125Yes
SciPy, original scale≈ 10.285714≈ 0.02650081125Yes
SciPy, scaled by 2^-511≈ 21.81818≈ 0.05611184474No

The reference uses numerator and denominator degrees of freedom of 2 and 4. These are properties of the reference calculation; f_oneway returns the statistic and p-value.

The reported environment is SciPy 1.18.1, NumPy 2.3.5, and Python 3.12.14 on Linux x86_64. The issue includes the reproducer, warning messages, and build configuration.

The overflow occurs when the weights are added

Welch’s ANOVA gives each group a weight equal to its sample size divided by its sample variance. In this example, each sample variance stays within the normal float64 range, and every individual weight stays finite. Their sum exceeds the largest value that float64 can represent.

That overflow affects the later calculation and produces the reported change in the statistic and p-value. This is the cause analysis in our report; it is separate from a maintainer’s confirmation or choice of repair.

The saved diagnostic script captured two RuntimeWarning: overflow encountered in reduce messages from the scaled f_oneway call with the warning filter set to always. The issue identifies these as captured warning categories and messages, with no recorded source filenames or line numbers. Default warning filtering may suppress a repeated message.

What this establishes

A p-value can remain between zero and one while changing a statistical decision because of an intermediate numerical failure. Checking a transformation that should preserve the result provides another way to expose that failure.

For Licklider, this case demonstrates the use of small reproducible examples and mathematical references to investigate scientific software. The scale was chosen deliberately for a boundary test. We have not estimated how often practical datasets encounter it, and this report does not validate a general repair.

Welch’s multi-group ANOVA and the two-group Welch t-test are different procedures. This report does not add Welch ANOVA verification to nomue; the current product limits remain the reference for supported use.

Read the upstream report and reproducer for the exact inputs, expected result, warnings, and environment.